HPCLab banner

Modern computer systems are becoming more diverse as power efficiency is increasingly a first priority for their operation. Today, the main drivers are AI’s massive compute and storage demands and the need for energy-efficient solutions. Heavy compute is offloaded to accelerators (GPU, FPGA, AI chips). Fast intelligent networks run functions in the fabric, bringing processing nearer data and easing load on compute nodes. Object stores replace limited POSIX file systems. The three pillars, (1) heterogeneous processors, (2) smart high‑performance networks, and (3) new storage, force continual redesign of algorithms and software. Future architectures may require new mathematical algorithms and implementation strategies. To keep development costs manageable across varied hardware, performance portability and productivity are essential. Early decisions on algorithm design and variants help stay aligned with evolving hardware, meet performance goals, and ensure sustainability.

The HPCLab develops algorithmic and software solutions for the efficient implementation of data-driven optimization and decision workflows on near-future processor architectures and storage technologies. Our focus is on optimization methods and simulation workflows designed and developed in the MODAL Labs EnergyLab, MedLab, MobilityLab, NanoLab, and SynLab.

Projects

In the third phase of the Research Campus MODAL, the HPCLab develops and refines solutions beneficial for application developers and system operators for leveraging heterogeneous systems, interconnects and object stores. In close cooperation with industrial and academic partners, including Cornelis Networks, NextSilicon, Intel and HPE, and international research partners, the lab addresses key challenges in heterogeneous computing, high-speed interconnects and fast object stores.

Performance and energy efficiency of heterogeneous systems.

We assess performance and portability on emerging processors using NextSilicon’s data‑flow design. We benchmark compute power and energy efficiency for MODAL algorithms, initiate the further development of the compiler/runtime system on NextSilicon, as well as performing optimisations for Nvidia, Intel, AMD GPUs (and optional FPGAs). In phase three, with NHR support, we focus on energy‑efficiency by gathering data, monitoring workflows, and running synthetic benchmarks of typical workloads. Energy analysis uses tools such as Energy‑Aware Runtime, perf, VTune, and LIKWID. HPCLab will share optimisation knowledge and deliver best‑practice guidelines to all MODAL teams.

Intelligent high-performance interconnects.

In partnership with Cornelis Networks, we are building a next‑generation Omni‑Path testbed to examine how network topology impacts AI/ML node performance and overall throughput, while also evaluating the benefits of programmable Host Fabric Interfaces (Smart NICs) on that platform; this research explores the synergy between configurable network hardware, development tools, and the resulting functional and performance improvements.

HPC object store DAOS.

Partnering with Cornelis Networks, HPE and Intel, we are boosting the next‑generation Omni‑Path interconnect to better support Intel DAOS and evaluating its native key‑value store for MODAL‑relevant application patterns.


Past Projects

IO500 Challenge (2024)

 

Together with partners Cornelis Networks and Intel, we successfully participated with NHR@ZIB's DAOS installation at the IO500 challenge and achieved the 3rd rank on the ten node production list among all competitors. Lise's DAOS achieved 65 GByte/s of bandwidth and 1,6 million I/O operations per second, resulting in twice the score of the system on rank four. This performance was achieved with 10 client nodes and 960 client processes. Within the full production IO500 list the same results placed Lise on rank five (read more).
Graph500 Challenge (2022)

 

Together with SynLab and NHR@ZIB, we successfully participated in the Graph500 challenge in the category of Breadth-First Search (BFS) with our Intel CLX partition using 122k cores and achieved the 9th rank among all competitors (see the November 2022 list here).

 

MPI for FPGA design

 

MPI for FPGAs. With the advent of high-level-synthesis-programmed and energy-efficient FPGAs, HPCLab explored how the emerging partitioned communication specification of the well-established communication standard MPI could be efficiently implemented with SYCL. A lightweight HLS-compatible implementation was successfully developed and achieved up to 85% of host-only communication bandwidth.

Publications

2026
A Flexible Open-Source Framework for FPGA-based Network-Attached Accelerators using SpinalHDL Architecture of Computing Systems - 39th International Conference, ARCS 2026, Mainz, Germany, March 24-26, 2026, Proceedings., 2026 (accepted for publication) Niklas Schelten, Steffen Christgau, Anton Schulte, Bettina Schnor, Hannes Signer, Benno Stabernack BibTeX
MODAL-HPCLab
Using FPGA-based Network-Attached Accelerators for Energy-Efficient AI Training in HPC Datacenters 2026 IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW), pp. 346-352, 2026 Niklas Schelten, Steffen Christgau, Merit Hutzler, Philipp Kreowsky, Marco De Lucia, Bettina Schnor, Hannes Signer, Johannes Spazier, Benno Stabernack, Serhii Yahdzhyiev BibTeX
DOI
MODAL-HPCLab
2025
On the Usability and Energy Efficiency of High-Level Synthesis for FPGA-based Network-Attached Accelerators 2025 IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW), pp. 886-895, 2025 Steffen Christgau, Everingham Dylan, Max Lübke, Marco De Lucia, Danny Puhan, Niklas Schelten, Bettina Schnor, Hannes Signer, Johannes Spazier, Benno Stabernack, Fritjof Steinert, Serhii Yahdzhyiev BibTeX
DOI
MODAL-HPCLab
2024
Gaining Cross-Platform Parallelism for HAL’s Molecular Dynamics Package using SYCL 29. PARS-Workshop 2023, pp. 66-77, Vol.36, Mitteilungen - Gesellschaft für Informatik e.V., Parallel-Algorithmen und Rechnerstrukturen, 2024 Viktor Skoblin, Felix Höfling, Steffen Christgau BibTeX
arXiv
URN
MODAL-HPCLab
2023
Enabling Communication with FPGA-based Network-attached Accelerators for HPC Workloads Proceedings of the SC'23 Workshops of The International Conference on High Performance Computing, Network, Storage, and Analysis, SC-W 2023, Denver, CO, USA, November 12-17, 2023, pp. 530-538, 2023 Steffen Christgau, Dylan Everingham, Florian Mikolajczak, Niklas Schelten, Bettina Schnor, Max Schroetter, Benno Stabernack, Fritjof Steinert BibTeX
DOI
MODAL-HPCLab
2022
A First Step towards Support for MPI Partitioned Communication on SYCL-programmed FPGAs IEEE/ACM International Workshop on Heterogeneous High-performance Reconfigurable Computing, H2RC@SC 2022, Dallas, TX, USA, November 13-18, 2022, pp. 9-17, 2022 Steffen Christgau, Marius Knaust, Thomas Steinke BibTeX
DOI
MODAL-HPCLab
An Early Scalability Study of Omni-Path Express Hamburg, 2022 Glenn Brook, Douglas Fuller, John Swinburne, Steffen Christgau, Matthias Läuter, Ronaldo Rodrigues Pelá, Lewin Stein, Tuma Christian, Thomas Steinke BibTeX
DOI
MODAL-HPCLab
Co-Design for Energy Efficient and Fast Genomic Search: Interleaved Bloom Filter on FPGA FPGA '22: Proceedings of the 2022 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays, pp. 180-189, 2022 Marius Knaust, Enrico Seiler, Knut Reinert, Thomas Steinke BibTeX
DOI
MODAL-HPCLab